Back

Journal of Pathology Informatics

Elsevier BV

Preprints posted in the last 90 days, ranked by how well they match Journal of Pathology Informatics's content profile, based on 15 papers previously published here. The average preprint has a 0.02% match score for this journal, so anything above that is already an above-average fit.

1
Vessel Spatial Analysis (VeSpA): a tool for whole slide image segmentation, morphometry, and QuPath extension.

Grion, G.; Hussain, R.; Colella, F. E.; Roufail, K.; Uccella, S.; Frapolli, R.; Matteo, C.; Mintemur, O.; Pennati, F.; Renne, S. L.

2026-06-20 pathology 10.64898/2026.06.15.732366 medRxiv
Top 0.1%
23.3%
Show abstract

Quantifying vascular architecture in histological whole slide images is needed to study tissue organisation, tumour microenvironment biology, and diseaseassociated vascular remodelling. However, vessel analysis in routine immunohistochemistry remains challenging. Available workflows are often manual, require programming expertise, or lack direct integration with digital pathology platforms. We developed VeSpA (Vessel Spatial Analysis), an open-source pipeline and QuPath extension for automated vessel segmentation and morphometric quantification in CD31-stained whole slide images. VeSpA combines configurable signal extraction, using CMYK Yellow channel extraction by default and optional DAB stain deconvolution for H-DAB images, with automatic or percentile-based thresholding, morphological refinement, contour filtering, and lumen filling to generate vessel masks from standard DAB-stained sections. The QuPath extension includes a graphical interface for selecting annotations, TMA cores, or whole images, configuring segmentation parameters, running the Python backend, and importing vessel objects directly into the QuPath hierarchy. For each detected vessel, VeSpA extracts area, major axis length, minor axis length, eccentricity, centroid, and orientation, while also appending summary measurements to parent annotations and TMA cores. Validation against independent pathologist annotations showed that VeSpA achieved segmentation performance close to inter-rater agreement and outperformed yellow channel prompt-based SAM and zero-shot YOLOv8-seg on overlap-based metrics in the tested dataset. VeSpA integrates vessel segmentation, morphometric feature extraction, and QuPath-based visualisation into a single reproducible workflow for vascular quantification in computational pathology and spatial analysis of histological tissue architecture.

2
Artificial intelligence-assisted ganglion cell detection in Hirschsprung's disease: A comparative evaluation of two deep learning approaches

Wang, E.; Grenier, K.; Savadjiev, P.; Poenaru, D. D.

2026-06-12 pathology 10.64898/2026.06.11.26354826 medRxiv
Top 0.1%
18.9%
Show abstract

Background. Definitive diagnosis of Hirschsprung's disease (HD) requires pathological identification of enteric ganglion cells. This process is time-consuming and subject to inter-observer variability. Artificial intelligence (AI) tools have the potential to standardize and accelerate this workflow, but no study has determined which AI approach best serves intraoperative HD pathology diagnostics. Method. This study compared the U-Net and You Only Look Once version 26 (YOLO26) frameworks for ganglion cell detection using a single-centre retrospective dataset of 54 whole-slide images (WSIs) from rectal biopsies. WSIs were tiled into 397,731 image patches (128x128 pixels), further partitioned into training (70%), validation (15%), and testing (15%) sets. Models were evaluated on tile- and patient-level diagnostic metrics and processing latency. Results. The U-Net achieved a tile-level sensitivity of 82.9%, showing no statistically significant difference compared to YOLO26 (79.1%; p = 0.097). However, YOLO26 demonstrated a statistically significant advantage in tile-level specificity (96.1% vs. 93.9%; p < 0.001) and reduced mean inference latency (7.64 ms vs. 11.57 ms/tile). At the patient level, both models achieved 100% diagnostic sensitivity. Despite low patient-level specificity (0.0% U-Net; 11.8% YOLO26), the tissue-level diagnostic burden of false positives was 6.00% for U-Net and 3.50% for YOLO26. Conclusion. The U-Net is preferred when nominal gains in sensitivity are prioritized, while the YOLO26 is an alternative that optimizes efficiency and false positive suppression. Both models serve as robust screening filters to augment the pathologist's workflow and should be selected based on workflow requirements. Prospective validation on larger, multi-centre datasets is required before clinical implementation.

3
Cross-country generalizability of foundation models for cervical cancer screenings on H&E whole slide images

Paulikat, M.; Bosch, C.; Aswolinskiy, W.; Caixeta Borges, I.; Nauschuette, L.; Aichmueller, C.; Schmidt, D.; Bussmann, H.; Kalteis, S.; Zapukhlyak, M.; von Knebel Doeberitz, M.; Kloor, M.

2026-07-23 pathology 10.64898/2026.07.22.26358575 medRxiv
Top 0.1%
15.2%
Show abstract

Accurate grading of cervical biopsies on Hematoxylin and Eosin (H&E) stained whole slide images (WSIs) is essential for distinguishing high grade lesions from low grade changes, yet this process is subject to considerable inter-observer variability. In this study, we evaluate a foundation model-based multiple instance learning (MIL) pipeline for binary high-grade squamous intraepithelial lesion (HSIL) detection on H&E stained WSIs. We benchmark our Athena foundation model against four state-of-the-art pathology foundation models: H-optimus-0, Hibou-L, Midnight-12k and Virchow, across datasets from five different countries: Portugal, Cambodia, Germany, Poland and Scotland. Athena achieved the highest mean area under the curve (AUC) (0.931) with the lowest cross-country variability (STD = 0.022). Furthermore, we compared the model's diagnostic performance to that of trained pathologists on a dataset with p16-confirmed ground truth. Our model improved sensitivity from 84% to 95% while maintaining comparable specificity (85% vs. 84%). Failure analysis revealed that the model's errors were concentrated at the diagnostic boundary between low-grade and high-grade lesions, whereas pathologists' errors spanned a broader range of misclassifications. These findings show the potential of foundation models for cervical cancer screenings worldwide.

4
An Agentic, No Code Artificial Intelligence Workflow for Developing and Externally Validating a Thyroid Nodule Ultrasound Malignancy Classifier

Thomas, J.; Pozdeyev, N.

2026-06-26 endocrinology 10.64898/2026.06.23.26356395 medRxiv
Top 0.1%
13.1%
Show abstract

Convolutional neural networks (CNNs) can classify thyroid nodules on ultrasound, yet published models are seldom available for independent testing, require machine learning expertise to develop and deploy, and are validated mostly on papillary thyroid carcinoma. Objective. To test whether an autonomous (agentic), no code artificial intelligence (AI) agent can develop a calibrated thyroid-nodule malignancy classifier, and to validate it internally and on an external cohort spanning multiple cancer histologies. Methods. This is a retrospective, computational diagnostic study with prespecified endpoints. A no code agent (Hugging Face ML Intern) autonomously reviewed data, selected and trained the model and calibrated probabilities, using the open source TN5000 dataset (3500 training, 500 validation, and 1000 test images). The trained ResNet 18 model was externally validated on 232 nodules from the University of Colorado, including follicular, medullary, oncocytic, and follicular variant of papillary carcinomas. Results. On the internal test set, an agentic AI model achieved AUROC 0.94 (95% CI, 0.920 - 0.953), sensitivity 0.90, and specificity 0.80. On external validation, agentic AI model achieved an AUROC of 0.90 (95% CI, 0.850 - 0.936), sensitivity of 0.92, and specificity of 0.68, negative predictive value of 0.96, and positive predictive value of 0.52, exceeding the performance of a previously published classifier on the same cohort (AUROC of 0.83). Conclusions. An agentic, no code AI workflow produced a calibrated, externally validated thyroid nodule classifier, supporting accessible, reproducible, and independently testable medical AI development. Prospective validation and local recalibration are required before clinical use.

5
PRECISE: Benchmarking digital pathology with expert-annotated contiguous IHC-H&E serial prostate sections

Calapaqui Teran, A. K.; Gonzalez Bernad, A. A.; Cobo Cano, M.; Sanchez Magdaleno, L.; Marcos Gonzalez, S.; Delgado Bolton, R. C.; Moustafa Calvo, J.; Gomez Roman, J. J.; Lara, L.

2026-07-22 pathology 10.64898/2026.07.21.26358559 medRxiv
Top 0.1%
6.0%
Show abstract

We present PRECISE (PRostate Expert-annotated Contiguous IHC-H\&E Serial sEctions), a hybrid histopathology dataset of paired hematoxylin and eosin (H\&E) and immunohistochemistry (IHC) whole-slide images (WSIs), comprising 37 prostate core needle biopsies from 25 patients, each with matched H\&E and CKAPM+racemase staining. To the best of our knowledge, this is the first publicly available dataset offering spatially harmonized, pixel-level expert annotations across both staining modalities in prostate biopsy WSIs - directly mirroring the two-stage (H\&E-then-IHC) clinical diagnostic workflow used to resolve morphological uncertainty, restricted to cases in which that workflow reached diagnostic consensus. The dataset contains 24,387 annotations spanning seven diagnostically critical classes: malignant glands, benign glands, stromal tissue, intraductal carcinoma (IDC-P), high-grade prostatic intraepithelial neoplasia (HGPIN), atypical intraductal proliferation (AIP), and tissue artifacts. Unlike existing resources, which focus on binary tumor classification or lack IHC pairing, this dataset captures the full morphological spectrum encountered in routine prostate pathology, including rare precursor lesions and confounding entities underrepresented in current benchmarks. Annotations were validated through a structured three-stage consensus by two expert uropathologists, with IHC serving as biological ground truth for boundary definition. PRECISE is designed as a robust benchmark for multimodal semantic segmentation and self-supervised learning, and is openly released to promote reproducible research and accelerate AI-assisted diagnosis in prostate cancer.

6
Cancer-Tissue Fraction as a Scanner-Robust Triage Signal for Automated Gleason Grading of Prostate Biopsies: External Validation Across a Middle Eastern Cohort

Ebbert, J. L.; Perry, A.; Szymanski, J.; Della Corte, D.

2026-07-29 pathology 10.64898/2026.07.28.26359127 medRxiv
Top 0.1%
5.6%
Show abstract

Background: Deep-learning systems for Gleason grading are developed almost entirely on high-end clinical scanners and on cohorts from a small number of Western institutions, yet deployment increasingly involves other devices and other populations. These two distribution shifts, device and population, are rarely tested together on the same physical slides. The PAR dataset, from Erbil, Iraq, digitizes each biopsy on three scanners and provides three distinct pathologist grades, so it permits both tests at once on a Middle Eastern cohort. A concurrent study by the dataset originators validated a task-specific model and two foundation models on PAR; we complement it by testing an inde-pendently developed detect-then-grade pipeline and by separating scanner effects on detection from scanner effects on grading. Methods: We applied one fixed model de-veloped on North American and European material to all 1017 whole-slide images (339 slides from 185 patients, three scanners; 49.6% clinically significant cancer) with no scanner-specific or population-specific tuning. We measured cancer detection (area under the ROC curve of the predicted cancer-tissue fraction), all-slide ISUP agreement of the deployed detect-then-grade pipeline (quadratic-weighted kappa, QWK), and grading agreement on pathologist-confirmed cancers, at the slide level and, because a case carries up to two slides, at the patient level. The reference reader was S.A.; thresholds and operating points were cross-validated leave-one-out; scanners were compared by paired within-biopsy bootstrap and confidence intervals confirmed by patient-cluster bootstrap. Results: Detection was statistically equivalent across scanners (AUC 0.987 to 0.991; paired differences at most 0.003) and transferred to this non-Western cohort with no per-population tuning. At a 95% sensitivity operating point the deployed pipeline reached cross-validated all-slide QWK of 0.86, 0.81, and 0.86 (Grundium, Hamamatsu, Leica), matching the inter-pathologist ceiling of 0.81, against 0.23 to 0.62 for the ungated model. Grading of confirmed cancers was scanner dependent: the compact Grundium (0.63) did not differ from the clinical Leica (0.67; paired difference 0.04, 95% CI -0.03 to 0.11), while both exceeded Hamamatsu (0.44). Results held at the patient level, with grading somewhat lower for two scanners; the two slides of a case disagreed in grade in 43% of cases, and patient clustering did not widen the intervals. Conclusions: Can-cer-tissue fraction is a triage signal robust across scanner and transferable to an un-derrepresented population for detection, while grading is the scanner-sensitive step. Prostate grading models should be deployed as a detect-then-grade pipeline, with grading validated per device and confirmed on the local population.

7
A Multicenter Swedish Histopathology Image Dataset Of Pediatric Central Nervous System Tumors

NYMAN, P.; Tampu, I. E.; Shamikh, A.; Prochazka, G.; Blystad, i.; Basmaci, E.; Diaz de Stahl, T.; Augustsson, P.; Zielinska-Chomej, K.; Cao, D.; von Salome, J.; Ardalan, A.; Somarajan, P. R.; Ljungman, G.; Lundberg, P.; Sandgren, J.; Haj-Hosseini, N.

2026-06-16 pathology 10.64898/2026.06.15.26355523 medRxiv
Top 0.1%
5.5%
Show abstract

Refined detection methods, more detailed tumor characterization, and adequate distinction between different pediatric tumor subtypes are necessary to improve diagnosis and treatment, enable precision medicine, and advance patient prognosis. However, the application of computational approaches to pediatric brain tumors remains limited, largely due to the lack of accessible datasets. To address part of this gap, we provide whole slide images (WSIs) of hematoxylin and eosin (H&E)-stained tissue sections from all pediatric central nervous system (CNS) samples collected in Sweden between 2013 and 2023. These data represent a population-based national cohort encompassing all six pediatric oncology centers in Sweden and are available through the Swedish Childhood Tumor Biobank (BTB). The dataset includes 1,446 WSIs of sufficient image quality with confirmed CNS tumor diagnoses, derived from 537 unique subjects (562 cases). In addition, diagnosticrelevant clinical information is included. Corresponding whole-genome sequencing (WGS), wholetranscriptome sequencing (WTS), and methylation array data are available for most tumor samples through separate resources. This H&E dataset has been specifically curated to support artificial intelligence-based analyses, while also serving broader applications in medical research and education. When combined with matched molecular data, it provides a valuable resource for advancing multimodal and precision diagnostic approaches in the pediatric population. Refined detection methods, more detailed tumor mapping and adequate distinction between different subtypes of pediatric tumors are necessary to improve treatment, enable precision medicine and improve patient prognosis. Application of computational algorithms for pediatric brain tumors is very limited mainly due to the unavailability of pediatric histology brain tumor data sets. To enable the development of AI models comprehensive datasets covering a wide range of pediatric brain tumors are needed.

8
Three multimodal large language models fail at clinically actionable breast pathology in three different directions

Kang, Y.-J.; Jun, S.-Y.; Kim, S.

2026-06-22 pathology 10.64898/2026.06.18.26355928 medRxiv
Top 0.1%
4.8%
Show abstract

Background. Breast cancer treatment depends on histopathological features, such as grade and receptor-defined subtype; however, specialist pathologist access is constrained when the workforce is limited. Commercial multimodal large language models (MLLMs) accept hematoxylin and eosin (H&E) image tiles through paid interfaces without local hardware or fine-tuning. However, prior pathology evaluations addressed only coarse tasks. Whether they reach treatment-determining accuracy and whether vendors agree remain unclear. Methods. We aimed to evaluate three vendor-designated flagship MLLMs (Claude Sonnet 4.6, Gemini 2.5 Pro, GPT-5.5) in 427 invasive breast cancer cases. Each case went to all three with identical H&E tiles and prompts, and the subtype was inferred in the second call. The reference was an institutional sign-out report of an immunohistochemistry-derived subtype. We calculated the concordance, sensitivity, specificity, Cohen's kappa, and pairwise McNemar and Bowker tests. Findings. Claude ranked highest by raw histologic-type concordance but lowest by kappa, classifying all 23 lobular and seven micropapillary carcinomas as invasive breast carcinoma of no special type. The models anchored the Nottingham grade to three modal grades. None of the models reliably identified human epidermal growth factor receptor 2-positive disease. The failure direction was vendor-specific: Claude and GPT-5.5 were under-detected, whereas Gemini was over-called. Twelve prompt variants (4,056 calls) did not recover sensitivity. Interpretation. No current commercial MLLM reaches deployment-ready accuracy for any treatment-determining feature of breast pathology. As each vendor fails in its own fixed direction, changing vendors alters the type of error rather than removing it; therefore, the value of these models is assistive rather than autonomous. At USD 0.20-0.50 per case, they may serve as supervised draft generators that leave the diagnosis with the pathologist.

9
Intra-slide calibration technology improves immunohistochemical harmonization within and between anatomic pathology laboratories

Fernandes, G. M. d. M.; Wang, W.; Parwani, A.; Ahmadian, S. S.; Alves, M. J.; Philips, J. J.; Otero, J. J.

2026-06-08 bioinformatics 10.64898/2026.06.04.730099 medRxiv
Top 0.1%
3.5%
Show abstract

The reproducibility of immunohistochemistry in tumor tissue analysis across reference labs remains a persistent challenge. We tested the extent to which an intra-slide calibration technology mitigated discprepencies in inter-laboratory assays of p53 immunohistochemical (IHC) reactions in brain biopsies of glioblastoma (GB), IDH-wildtype. Intra-slide calibration technologies apply a 0-100% concentration scale incorporating primary surrogate and secondary antibodies to generate a standardized curve for DAB precipitation. IHC from GB samples was performed independently by pathology departments from two different hospital laboratories and were digitalized at 40x magnification using Aperio Image Scope software. Feature extraction, including intensity and texture parameters was performed using the EBImage package in R, followed by UMAP dimensionality reduction and DBSCAN clustering analysis. Our results show significant differences in intensity and texture clustering patterns between laboratory tissue samples and intra-slide calibration technology ruler caused by the different laboratories. Intra-slide calibration technology coupled with polynomial regression analysis improved ~90% the data harmonization. Our findings demonstrate a key role for computational pathology using intra-slide calibration technology to enable intra-laboratory consistency and inter-laboratory reproducibility. These advances strengthen the reproducibility of diagnostic assessments and support more objective, data-driven decision-making in neuro-oncology.

10
Homologous recombination deficiency prediction from whole slide images using label refinement and foundation-model benchmarking in ovarian cancer

Shah, N. A.; Sarwar, M.; Ullah, E.

2026-06-30 pathology 10.64898/2026.06.25.734452 medRxiv
Top 0.1%
3.2%
Show abstract

Background: Homologous recombination deficiency (HRD) is clinically imperative in high-grade serous ovarian carcinoma (HGSOC), particularly because of its association with platinum sensitivity and benefit from poly(ADP-ribose) polymerase inhibitor (PARPi) therapy. However, public datasets rarely contain a complete combination of diagnostic haematoxylin and eosin (H&E) whole-slide images (WSIs), validated clinical HRD assay results, genomic scar scores, BRCA1 promoter methylation data, and treatment-response outcomes. This creates a major barrier for computational pathology studies seeking to develop clinically interpretable models of HRD or PARPi response from routine histology. Objective: We performed an exploratory, leakage-controlled computational pathology benchmarking study to evaluate whether H&E WSIs from TCGA-OV contain a measurable morphology-linked signal associated with research-grade molecular HRD labels, and whether label refinement and pathology foundation-model embeddings alter predictive performance. Methods: We assembled a frozen-primary TCGA-OV WSI cohort comprising 717 tissue-section/biospecimen slides from 316 patients. Diagnostic FFPE DX slides were excluded from model selection because of complete patient overlap with the frozen-primary cohort. Two HRD labels were evaluated: an initial mutation-only molecular label based on BRCA/HR-gene mutation evidence, and a refined methylation-enhanced molecular label that additionally incorporated BRCA1 promoter methylation. Feature extraction was performed using ResNet50, UNI, CONCH, Virchow2, Phikon-v2, and UNI2-h encoders. Patient-level attention-based multiple instance learning (ABMIL) was used with patient-as-bag modelling. Evaluation used patient-level grouped 5-fold x 5-repeat stratified cross-validation, with 25 folds total, bootstrap confidence intervals, and patient-level leakage control. Results: The initial mutation-only label classified 78 patients as positive and 238 as negative. The refined methylation-enhanced label recovered 33 additional positives, resulting in 111 positive and 205 negative patients. Patient-level ABMIL using UNI2-h features achieved the strongest performance for the refined label, with AUROC 0.634 (95% CI 0.571-0.698), AUPRC 0.468 (95% CI 0.390-0.562), balanced accuracy 0.597, sensitivity 0.532, specificity 0.663, F1 score 0.494, and Brier score 0.233. The calibrated threshold was 0.512, yielding TN=136, FP=69, FN=52, and TP=59. Comparative models showed lower discrimination, including UNI2-h with the initial label (AUROC 0.628), Phikon-v2 refined (0.582), Virchow2 refined (0.582), CONCH initial (0.587), ResNet50 refined (0.570), and clinical baselines (AUROC 0.54-0.57). Conclusions: TCGA-OV H&E WSIs contain a modest but reproducible morphology-linked signal associated with research-grade molecular HRD status. However, the AUROC around 0.63, absence of clinical HRD assay labels, lack of genomic scar endpoints in the implemented workflow, and absence of PARPi/platinum response targets prevent clinical interpretation. This study should be interpreted as a proof-of-concept benchmarking framework and methodological foundation for future H&E-based predictive modelling in clinically curated PARPi response cohorts.

11
Development of a multiplex immunofluorescence panel to study heterogenous cancer-associated fibroblast subtypes with spatial resolution

Burley, A.; Silveira, T.; James, N.; Salto-Tellez, M.; Wilkins, A. C.

2026-07-01 pathology 10.64898/2026.06.26.734718 medRxiv
Top 0.1%
2.7%
Show abstract

Background: Single cell RNA sequencing provides a wealth of information to explore the complexities of the tumour microenvironment, but crucially the spatial topology of the tumour is lost and studying cellular interactions is limited. Spatial transcriptomics aims to address this however the technique remains cost prohibitive for the generation of data from meaningfully-sized clinical cohorts. In contrast, spatial proteomic profiling with multiplex immunofluorescence, preserves spatial interactions, is relatively cost accessible, and is scalable for large clinical cohorts to address powerful translational questions. Whilst multiplex approaches have advanced in recent years, we note that cancer-associated fibroblasts (CAFs) have been explored in less detail, potentially due to difficulties associated with CAF heterogeneity and the diversity of markers used to define them. Methods: We designed, optimised, and validated a multiplex immunofluorescence panel that combines four frequently used CAF markers; alpha smooth muscle actin (aSMA), fibroblast activation protein (FAP), podoplanin (PDPN) and platelet-derived growth factor receptor alpha (PDGFRa) with CD8 and pan-cytokeratin. Here we share our methodology and the practical considerations taken to inform the final panel design. We also highlight the benefits of robust optimisation experiments.

12
Multi-model Segmentation and Morphometric Quantification of Cerebral Amyloid Angiopathy in Alzheimer's Disease Whole Slide Histopathology Images

Tahmasebidehkordi, H.; Bahramy, A.; Julian, D. R.; Cohen, J. A.; Neal, M.; Bumgardner, C.; Nelson, P. T.; Pearce, T. M.; Kofler, J.

2026-07-21 pathology 10.64898/2026.07.16.739032 medRxiv
Top 0.1%
2.4%
Show abstract

IntroductionCerebral amyloid angiopathy (CAA) is characterized by amyloid-beta deposition in cortical and leptomeningeal vessels and associated with cognitive impairment and hemorrhage. Current neuropathological assessments rely on semiquantitative grading and lack vessel-level resolution and scalability. Existing computational pathology approaches also fail to capture individual vessel morphology and spatial amyloid distribution across whole-slide images (WSIs). To address this gap, we developed a deep learning framework for reproducible, quantitative analysis of CAA in WSIs. MethodsWe analyzed 20 postmortem brain tissue sections from the frontal (n = 10) and occipital cortices (n = 10) of 10 individuals with Alzheimers disease pathology obtained from the University of Pittsburgh Alzheimers Disease Research Center, which served as the internal development cohort. An independent external cohort consisted of 10 sections (5 frontal and 5 occipital samples) from 5 individuals obtained from the University of Kentucky Alzheimers Disease Research Center. We trained and compared three semantic segmentation architectures, a standard U-Net, a dual-attention residual U-Net (DA-ResUNet), and a Swin Transformer-based U-Net (Swin-UNet), using the internal development cohort with slide-level five-fold cross-validation. All models were evaluated on the independent external cohort to assess generalization under domain shift. Based on segmentation performance and computational efficiency, we selected one architecture to generate whole-slide composite segmentation masks for vessel walls, amyloid deposits, and tissue compartments. These masks were subsequently used for deterministic vessel detection, morphometric measurements, and quantification of vascular and perivascular amyloid features through post-processing analysis. ResultsAll three architectures achieved high segmentation accuracy on the internal cohort, with Dice scores above 90% across vessel walls, amyloid deposits, gray matter, and leptomeninges. The Swin-UNet showed marginally higher performance for vessel segmentation, whereas the DA-ResUNet provided more balanced accuracy and computational efficiency and was selected for downstream analysis. External cohort evaluation demonstrated robust generalization, with attention-enhanced models outperforming the standard U-Net under domain shift. Using the selected model, the pipeline reliably detected valid vessels, excluded non-vascular artifacts, and enabled deterministic extraction of vessel morphometry, vascular and perivascular amyloid burden, and identification of circumferential CAA involvement at the vessel level. DiscussionThis framework provides a scalable, interpretable solution for vessel-level CAA analysis, supporting robust geometric and spatial characterization of cerebrovascular pathology and enabling future integration with clinical and genetic studies. Beyond CAA, the modular design allows extension to other vascular pathologies, including arteriolosclerosis, in WSIs, facilitating broader investigation of cerebrovascular disease mechanisms.

13
Ferumoxytol dynamic contrast-enhanced MRI for in vivo longitudinal cotyledon perfusion assessment with pathology correlation in a rhesus macaque thrombotic injury model

Liu, R.-Y.; Keding, L. T.; Edmondson, R.; Vazquez, J.; Antony, K. M.; Johnson, K. M.; Shah, D. M.; Golos, T. G.; Stanic, A. K.; Wieben, O.

2026-08-10 pathology 10.64898/2026.08.04.742075 medRxiv
Top 0.1%
2.2%
Show abstract

IntroductionWhile placental perfusion and pathology jointly affect pregnancy outcomes, cotyledon-specific perfusion across gestation and its correlation with local injury is not yet well understood. Ferumoxytol dynamic contrast-enhanced magnetic resonance imaging (DCE-MRI) offers a promising way to noninvasively identify cotyledons across gestation and quantify longitudinal cotyledon-specific perfusion changes. Additionally, intraplacental injection of bioactive fibrin sealant allows us to model thrombotic placental injury and further assess cotyledon-level relationships between perfusion and significant injury. MethodsPregnant rhesus macaques (N=13) received intrauterine saline or fibrin sealant injections at gestational day (GD) [~]101 and underwent ferumoxytol DCE-MRI at GDs [~]100, 115, and 145. Placental perfusion domains derived from contrast arrival time were segmented at each imaging time point and matched to cotyledons identified following tissue collection by cesarean section, with cotyledon perfusion quantified longitudinally and correlated with cotyledon-specific quantitative histopathology. ResultsAll pregnancies were successfully carried to term. Fibrin sealant injections induced significantly higher levels of placental pathology compared to saline controls. MRI-derived perfusion domains were largely consistent across gestation and showed predominantly one-to-one correspondence with term cotyledons, with successful perfusion-pathology pairing achieved in 153 cotyledons. Longitudinal cotyledon perfusion changes showed significant positive correlations with villous agglutination injuries. ConclusionsFeasibility of noninvasively tracking placental cotyledon perfusion using ferumoxytol DCE-MRI was demonstrated, and the efficacy of the rhesus macaque thrombotic injury model was confirmed. The positive perfusion-pathology correlations suggested intrinsic placental regulatory mechanisms and functional plasticity. This new framework is promising for future translational studies and validation of ex vivo cotyledon perfusion models. HighlightsO_LILongitudinal tracking of placental perfusion domains with ferumoxytol MRI C_LIO_LISuccessful matching of cotyledons and MRI-derived perfusion domains C_LIO_LIConfirmed thrombotic injury-model induced cotyledon pathology C_LIO_LIMaternal perfusion compensation in presence of villous pathology C_LI

14
Label-Free Threshold Selection for Out-of-Distribution Detection in Liver CT Segmentation

Nielsen, M.; Castelo, A.; Altaie, M.; Bennett, J.; Anthony, A.; Siddiqi, N. S.; Gupta, A. C.; Brock, K. K.; Woodland, M.

2026-08-24 radiology and imaging 10.64898/2026.08.20.26360809 medRxiv
Top 0.1%
2.2%
Show abstract

Reliable clinical deployment of automated liver segmentation requires mechanisms for detecting failures in rare and previously unseen scenarios. Achieving this goal requires an appropriately calibrated threshold that converts an out-of-distribution (OOD) score into a failure prediction. However, threshold calibration typically relies on expert-labeled failures, creating a substantial annotation burden when failures are rare. Building upon our prior work, which uses Pairwise Surface DSC scores as indicators of segmentation quality, we propose a label-free framework for calibrating OOD score thresholds. First, we fitted a log-t distribution to Pairwise Surface DSC scores from a validation set of 400 internal scans to approximate an in-distribution score distribution. New segmentations were assigned significance scores based on their extremity under this fitted distribution and categorized into Low, Medium, and High Risk review groups using statistically principled cutoffs of 0.25 and 0.05. The fitted log-t distribution provided a strong fit to the observed scores and remained robust to moderate contamination by OOD cases. On an independent test set of 500 internal and external scans, the combined Medium and High Risk categories achieved 100% sensitivity and 79% specificity, whereas the High Risk category alone achieved 78% sensitivity and 96% specificity. These results indicate that clinically meaningful failure detection can be derived from unlabeled data. Our code is available at https://github.com/marshalln7/Label_Free_OOD_Threshold_Selection.

15
Histological triage of early-stage mycosis fungoides using a weakly supervised deep learning-based model: a multicentre, external validation, and clinical utility study

Doeleman, T.; Brussee, S.; Valkema, P.; Kempf, W.; Vermeer, M.; Kers, J.; Wynaendts, L.; Kerckhoffs, K.; de Jonge, M.; Nguyen, A.; Peters, E.; Wobser, M.; Rauert-Wunderlich, H.; Rosenwald, A.; Stadler, R.; Jansen, P.; Battistella, M.; Roccuzzo, G.; Quaglino, P.; Schrader, A.

2026-07-28 pathology 10.64898/2026.07.27.26359009 medRxiv
Top 0.1%
2.1%
Show abstract

Background Histological diagnosis of early-stage mycosis fungoides (MF) is hindered by profound overlap with benign inflammatory dermatoses (BIDs), leading to diagnostic delays and extensive ancillary testing. We developed MIMIC (Multiple Instance-learning for Identification of Mycosis fungoides In Cutaneous biopsies), a weakly supervised deep learning model designed as a triage tool at initial H&E whole slide image (WSI) review to distinguish classic patch and plaque stage MF from BIDs. We externally validated the model and evaluated its clinical utility. Methods In this retrospective multicentre study, we trained a base model using weakly supervised attention based multiple instance learning on 3,339 WSIs from two Dutch centres. Crucially, all MF training labels were derived from a deeply phenotyped national cohort featuring strict multidisciplinary expert panel consensus diagnoses (the clinical gold standard). Transportability was evaluated on 371 WSIs from four independent European centres. A blinded reader study on 171 WSIs compared morphology only performance of MIMIC with 11 (dermato-)pathologists. We then retrained an updated model on all retrospective multicentre data and assessed clinical utility in a strictly held out, consecutive Utrecht cohort (2022-2023; 486 accessions, 863 WSIs). Primary analysis focused on classic MF versus BIDs (453 accessions). Decision curve analysis, using Platt scaled probabilities to correct for spectrum bias, evaluated net benefit at a prespecified, safety oriented threshold of 0.04. Findings The base model showed good multicentre transportability (mean centre specific AUROC 0.91; pooled AUROC 0.84). In the reader study, MIMIC achieved an AUROC of 0.87, exceeding the mean pathologist AUROC (0.79) and the best individual reader (0.83). In the consecutive MF versus BID cohort, the updated model achieved an AUROC of 0.87 (95% CI 0.81-0.92). At the 0.04 threshold, sensitivity was 97.8% (44/45 MF cases) and specificity 50.2%, reducing unnecessary ancillary workups by 39.9 per 100 screening cases versus a test all strategy. Interpretation By identifying nearly half of BIDs as low risk while preserving near complete sensitivity for classic early stage MF in a European digital pathology workflow, this unimodal H&E approach offers a scalable digital solution to reduce defensive ancillary testing and accelerate the diagnostic journey for patients with MF. Further validation is needed in non European centres and in populations with darker skin phototypes.

16
Serial Immunohistochemistry for High-Dimensional Single-Cell Spatial Analysis of Human Kidney Biopsies

Yang, X.; Marlin, M. C.; Celia, A. I.; Lee, C.-Y.; Cammarata-Mouchtouris, A.; Stephens, T.; Haddad, M.; Bradshaw, L.; Saksena, D.; Buyon, J.; Izmirly, P. M.; Putterman, C.; Kamen, D.; Petri, M.; Accelerating Medicines Partnership: RA/SLE Network, ; James, J. A.; Guthridge, J. M.; Fava, A.; Rosenberg, A. Z.

2026-08-12 pathology 10.64898/2026.08.06.743188 medRxiv
Top 0.1%
2.1%
Show abstract

BackgroundTraditional immunohistochemistry (IHC) with chromogen detection has limited multiplex capacity, detecting at most 4 protein markers per tissue section simultaneously, thereby restricting comprehensive spatial analysis of valuable human biopsies. We developed and validated a robust serial IHC (sIHC) staining method to detect multiple antigens on a single kidney biopsy slide, maximizing data yield for diagnosing and studying complex kidney diseases. MethodsFormalin-fixed, paraffin-embedded kidney biopsy sections were subjected to repeated IHC/imaging cycles with antibody removal using an optimized sodium dodecyl sulfate-glycerol buffer stripping protocol. Images were then co-registered, and analysis was performed using a variety of methodologies, including color deconvolution, cell segmentation, and spatial clustering. ResultsThis optimized sIHC method successfully detected up to 20 antigens on a single slide. Combining image analysis and artificial intelligence software, for example with HALO (Indica Labs), the assay assembles high-dimensional images and enables quantitative histology and single-cell spatial analysis. Using this advanced method, we were able to identify rare cell populations, such as double-negative T cells, that are challenging to detect conventionally. ConclusionWe have developed a validated, high-capacity sIHC protocol that uses standard IHC procedures with commercially available, clinically validated off-the-shelf antibodies. This method is a valuable, cost-effective tool for obtaining extensive, high-dimensional single-cell-resolved spatial data from limited pathology samples, such as a human kidney biopsy.

17
Vision Language Models Fail to Reliably Detect Acute Myeloid Leukemia in Bone Marrow Smears

Schulze, F.; Loeffler, C.; Radoynova, M.; Winter, S.; Roellig, C.; Sockel, K.; Kroschinsky, F.; Bornhaeuser, M.; Middeke, J. M.; Kather, J. N.; Eckardt, J.-N.; Ghaffari Laleh, N.

2026-08-22 hematology 10.64898/2026.08.19.26359329 medRxiv
Top 0.1%
1.8%
Show abstract

Hematologic diagnostics and especially cytomorphologic assessment are time-intensive and require high levels of expertise. Vision Language Models (VLM) show promise in medical image analysis in radiology and histopathology, while an evaluation on detecting acute myeloid leukemia (AML) is lacking. Our goal was to evaluate three Vision Language Models regarding their diagnostic accuracy and safety in clinical decision support in detecting AML from digitized bone marrow smears (BMS). Whole slide images were obtained from bone marrow smears of 50 AML patients and 50 bone marrow donors. Ten representative fields of view per sample were extracted manually. Three VLMs were used, two of which are considered generalist models (Qwen3.5-397B-A17B-FP8, GLM-4.6V-FP8), while the other one is a medically adapted model (Medgemma-27b-it). All models performed zero-shot analysis using two prompting strategies: First, a context-rich prompt requesting reporting of WHO/FAB diagnostic criteria in a structured manner, and secondly a minimal prompt without specific hematologic context. Overall diagnostic accuracy was poor for all models as they exhibited the overwhelming tendency to classify most samples as leukemic: With context-rich prompts, GLM4.6 identified 90% of leukemic samples while also labeling 92% of bone marrow donors as AML. The medical specialist model MedGemma-27b showed similar failure, misclassifying 86% of healthy donors and correctly detecting AML in only 66% of cases. Qwen3.5 performed best under detailed prompting, achieving a specificity of 0.26 and accuracy of 0.51. Accuracy of all models improved with context-free prompts (accuracies range 0.47-0.79), yet they still lacked the ability to correctly distinguish between leukemia and healthy bone marrow. Qwen3.5 was the only model to maintain meaningful specificity (0.64) and correctly identified 94% of AML, yielding an overall accuracy of 0.79. Morphologic feature-level agreement with human expert reports was poor across all models, indicating poor recognition of cell-level morphologies. This failure is likely driven by the fact that pathology imaging archives are vastly scraped during model training while hematological samples are not as widely available and therefore, hematology is an out-of-bounds use-case for these models, rendering them currently unsuitable for clinical decision support in hematology.

18
Harnessing Pathology Foundation Models to Accelerate Lymphoma Diagnosis Through Automated Immunohistochemistry Triage

Zhu, M.; Li, A.; Safa, I.; Galera, P.; Hazoglou, M.; Vanderbilt, C.; Kamali, A.; Goldgof, G.; Veeraraghavan, H.; Jiang, J.; Ardon, O.; Geneslaw, L.; Dogan, A.

2026-08-12 pathology 10.64898/2026.08.11.26360085 medRxiv
Top 0.1%
1.7%
Show abstract

Pathologic diagnoses of hematopoietic diseases require immunohistochemistry (IHC) stains selected by pathologists upon preview of H&E-stained slides. This multi-step workflow can delay diagnostic turnaround time by days. Hence, we developed the Hematopathology Automatic Triaging System (HATS), which automates IHC panel ordering directly from H&E whole-slide images using pretrained pathology foundation model representations combined with attention-based multiple-instance learning. After the most comprehensive evaluation of pathology foundation models for hematologic malignancy classification to date, encompassing seven publicly available models, we trained HATS on 4,996 whole-slide images from 1,607 patients spanning the ten most common lymphoma diagnostic categories. HATS achieves 84% case-level subtype classification accuracy (0.962 ROC-AUC), translating to 92% IHC panel ordering accuracy. In a blinded reader study, HATS outperforms practicing pathologists at predicting lymphoma subtypes from morphology alone (85% vs 65%). In an independent real-world validation of 230 clinical cases, after directing 7 cases with scant tissue for manual review, HATS-ordered IHC panels were sufficient for diagnosis in 72.6% of cases. By automating the triaging step while preserving full pathologist oversight, HATS offers a safe and practical entry point for clinical AI adoption in pathology.

19
Modeling Biomarker-Guided Avoidance of Radical Cystectomy: Costs and Outcomes

Sholklapper, T. N.; Li, M.; Srivastava, A.; Wagh, A.; Handorf, E.; Beck, J. R.; Abbosh, P.

2026-08-19 urology 10.64898/2026.08.17.26360582 medRxiv
Top 0.1%
1.6%
Show abstract

Importance There is growing interest to avoid radical cystectomy (RC) in patients with muscle-invasive bladder cancer (MIBC) who receive neoadjuvant chemotherapy and achieve pathological complete response (ypCR). To achieve this goal, molecular biomarkers will likely need to be used to enhance clinical staging given the limitations of evaluation by cystoscopy, cytology, and cross-sectional imaging. There are no studies evaluating whether safe RC avoidance (SaRCA) would be cost effective and what impacts it would have on quality of life (QoL) and survival. Objective This study models the potential economic, QoL, and survival costs/benefits of a ypCR biomarker as it relates to SaRCA using a decision analysis and Markov Model (MM). Methods/Materials A decision tree and MM was created to compare the expected costs of initial treatment, QoL, and survival under one strategy where all patients undergo RC after neoadjuvant treatment versus an alternative strategy where all patients would be subjected to the biomarker test with biomarker-positive patients (those with presumed residual disease) undergoing RC, while biomarker-negative patients (presumed complete responders) would undergo surveillance for up to 20 years. ypCR rates to neoadjuvant therapy, survival with and without RC, quality adjusted life years (QALY), and costs were abstracted from the literature. Test cost, sensitivity, and specificity were also abstracted from the literature for multiple clinical or liquid biopsy approaches. Results Broadly, SaRCA approaches are cost effective with the exception of systematic endoscopic evaluation (SEE). All testing approaches result in higher QALY and overall life expectancy compared to no testing. The cost of the test is offset by decreased usage of RC to realize a cost savings. These domains are further improved when cisplatin-based chemotherapy is replaced with emerging neoadjuvant therapies. Conclusions and relevance Modeling supports the development of accurate biomarker tests which can distinguish residual disease states to enable SaRCA. Such a biomarker could be used to avoid an expensive and risky operation, and unexpectedly would provide a survival benefit by reducing the number of perioperative mortalities in patients achieving ypCR. Development of an accurate biomarker-based test is likely to reduce cost and increase QoL and survival. An accurate biomarker test would have utility for patients, payers, hospitals, and physicians.

20
Analytical perturbation reveals hidden instability of biological phenotypes

Piorkowska, N. J.; Ostromecki, A.; Franik, G.; Bizon, A.

2026-07-16 endocrinology 10.64898/2026.07.13.26357916 medRxiv
Top 0.2%
1.5%
Show abstract

Background Unsupervised machine learning has become a cornerstone of computational phenotyping across clinical medicine, genomics, imaging, and multi-omics research. However, phenotype discovery relies on a sequence of analytical decisions - including missing-data handling, preprocessing, dimensionality reduction, clustering methodology, and stochastic initialization - that are rarely evaluated collectively. Although clustering stability has been extensively investigated, the robustness of complete analytical workflows remains largely unexplored. Results We developed an Analytical Perturbation Framework that systematically quantifies the robustness of phenotype discovery by perturbing complete unsupervised learning workflows rather than individual clustering algorithms. Using a real-world cohort of 1,286 women with polycystic ovary syndrome (PCOS), we generated 116 valid analytical pipelines comprising alternative preprocessing strategies, missing-data handling methods, dimensionality reduction approaches, clustering algorithms, and random initializations. Agreement between independently generated phenotype solutions was consistently low (median Adjusted Rand Index = 0.079), indicating substantial sensitivity of phenotype discovery to routine analytical decisions. Variance decomposition identified preprocessing as the largest contributor to phenotype instability (22.8%), followed by clustering methodology (14.6%), whereas stochastic initialization explained only 3.1% of the observed variability. At the patient level, most individuals exhibited reproducible phenotype assignments (median Patient Robustness Score = 0.719), although a substantial subgroup showed markedly lower assignment stability. Feature perturbation analyses identified follicle-stimulating hormone, anti-thyroglobulin antibodies, anti-thyroid peroxidase antibodies, total testosterone, luteinizing hormone, and androstenedione as the strongest contributors to computational robustness, rather than biological importance. Finally, phenotype solutions demonstrating greater computational robustness also exhibited greater biological coherence during independent validation.